Value in Health
○ Elsevier BV
Preprints posted in the last 7 days, ranked by how well they match Value in Health's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Leonhardt, C.; Birrer, D.; Stauffer, M. F.; Toti, J. M. A.; Gallagher, I. J.; Skipworth, R. J. E.; Laird, B.; Kuemmerli, C.
Show abstract
Importance Non-inferiority trials are becoming increasingly popular in abdominal surgery. The non- inferiority margin is critical in the interpretation and conclusion of these trials. Objective This systematic review aims to assess the methodological and reporting quality of non- inferiority randomized controlled trials in abdominal surgery. Evidence Review Non-inferiority trials were systematically identified by searching Ovid Medline, Embase and the CENTRAL databases from 2006 until December 2025. Randomized controlled trials in adult patients with any type of abdominal surgical intervention in at least one trial arm and a sample size greater than or equal to 100 were eligible for inclusion. The primary outcome was the definition of the non- inferiority margin. Secondary outcomes were the reporting of the non-inferiority margin, the robustness of its estimation, the uncertainty of the point estimate and the adequacy of conclusions. Findings A total of 11 045 trials were identified, of which 101 were eligible, enrolling 44 370 patients. Most trials provided a rationale for the non-inferiority design, while six (5.9%) trials did not. Previous literature was commonly used (n=56; 55.4%), but the non-inferiority margin was most often based on a clinical fixed margin or on historical comparison of the treatment and the active comparator. Based on the margin, investigators tolerated substantially worse outcomes of the treatment compared to the comparator. Conclusions were appropriate based on the confidence interval and the predefined non- inferiority margin in 88 (87.1%) of trials. The clinical judgement of the conclusion was overall adequate. Confidence interval estimations were reported in 16 (15.8%) of trials. Simulation studies were limited by the reporting quality. Conclusions and Relevance Clinical fixed margins are commonly used in abdominal surgery non-inferiority randomized controlled trials, however, substantial shortcomings in reporting limit the interpretability and reproduction of study findings. Based on the findings of this study, guidance on surgical- specific non-inferiority margin definitions is needed.
Green, J. L.; Davies, H.; Russell, D. A.
Show abstract
Background: The relative merits of infrainguinal bypass and primary major lower limb amputation (MLLA) for chronic limb-threatening ischaemia (CLTI) remain uncertain, and the baseline profiles of patients selected for each strategy are poorly described. Methods: A systematic review and meta-analysis were undertaken in accordance with PRISMA 2020 and prospectively registered (PROSPERO: CRD42022356094). MEDLINE, Embase, CENTRAL, and CINAHL were searched from inception to March 2025. Prospective studies of adults with CLTI undergoing primary infrainguinal bypass or primary MLLA were eligible. Mortality, major adverse cardiovascular events (MACE) and subsequent amputation outcomes were synthesised using random-effects meta-analysis of proportions. Baseline comorbidity profiles were also extracted. Results: Twenty-seven studies involving 6,576 patients were included: 5,779 underwent infrainguinal bypass and 797 underwent MLLA. After bypass, pooled mortality was 3.7% at 30 days (95% CI 2.8%-4.9%, I2 = 49.4%), 18.5% at 1 year (95% CI 15.6%-21.9%, I2 = 62.3%), and 54.3% at 5 years (95% CI 50.5%-58.0%, I2 = 0%). After MLLA, pooled mortality was 9.2% at 30 days (95% CI 4.1%-19.3%, I2 = 73.5%), 28.5% at 1 year (95% CI 13.3%-51.0, I2 = 70.8%), and 39.9% at 2 years (95% CI 0.3%-99.3, I2 = 90.5%), although longer-term estimates were limited by sparse data and marked heterogeneity. Thirty-day MACE was 6.5% (95% CI 4.3%-9.7, I2 = 63.5%) after bypass and 2.8% after MLLA (95% CI 0.1%-37.6%, I2 = 0%). Early subsequent major amputation after bypass occurred in 3.9% of patients (95% CI 2.0%-7.7%, I2 = 91.2%), rising to 16.2% at 1 year (95% CI 12.6%-20.5%, I2 = 82.0%) and 33.3% at 3 years (95% CI 20.1%-49.8%, I2 = 0%). Early re-amputation after MLLA occurred in 10.9% of patients (95% CI 4.5%-24.4%, I2 = 40.3%). Baseline comorbidity burden was high in both groups, with substantial heterogeneity across studies. Conclusions: CLTI carries a poor prognosis regardless of treatment strategy. Infrainguinal bypass is associated with lower early mortality and better early limb preservation than primary MLLA, but long-term survival remains poor and later limb failure is common. Primary MLLA is not a low-risk alternative. Better contemporary comparative evidence utilising modern causal inference approaches is needed to support individualised decision-making.
Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.
Show abstract
Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.
SHI, J.; Gu, Q.; Pan, J.; Yang, A.; Fan, M.
Show abstract
To evaluate the cost-utility and 5-year budget impact of first-line olaparib plus abiraterone versus abiraterone alone for metastatic castration-resistant prostate cancer (mCRPC) in China after the eleventh round of volume-based procurement (VBP). The intention-to-treat (ITT) population was assigned primary decision-analytic weight; the prespecified BRCA1/2-mutated (BRCAm) subgroup was a supporting analysis.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.
Show abstract
Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.
Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.
Show abstract
Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.
Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.
Show abstract
Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.
Show abstract
SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Mannava, S.; Ramkumar, V.; Murthy, G.
Show abstract
Introduction Hearing loss (HL) affects over 1{middle dot}5 billion people globally and India shares a disproportionately high burden including Disabling Hearing Loss (DHL). HL affects an Individual socio-economically, but there are limited studies on the broader societal economic consequences of HL in India.Methods Using Cost-of-Illness (COI) approach, we studied the societal economic burden of HL in India. This study uses epidemiological and macroeconomic data and modelling to estimate the loss of Gross National Income (GNI) due to HL and DHL across three economic pathways. Uncertainty is evaluated using deterministic and Probabilistic Sensitivity Analyses (PSA).Results The model estimates that there are in India, 289 million and 85{middle dot}9 million people with HL and DHL respectively. Direct Loss of GNI and Indirect Loss of GNI (Caregiver burden) are estimated as INR 4,648{middle dot}4 billion (USD 55{middle dot}6 billion) and INR 3,268 billion (USD 39 billion) respectively. The Loss of GNI due to Low Education amongst those with HL is estimated as INR 1,041{middle dot}9 billion (USD 12{middle dot}45 billion).Discussion Economic burden of HL is presented across three pathways with Direct Loss of GNI due to DHL being the greatest. It also presents age stratified caregiver economic burden. The findings of the study help in estimating similar cost pathways, advocacy, and policy decisions towards reducing HL prevalence in India and LMICs. This study also highlights the need for India specific estimations related to the HL attributable low education, state-wise disaggregates, and prevalence studies. Funding This study has not received any funding.
Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.
Show abstract
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.
Pinedo-Torres, I.; Taype-Rondan, A.; Vera-Luza, A. A.; Zegarra-Lizana, P. A.; Rojas-Vilca, J. L.; Yovera-Aldana, M.
Show abstract
Objective. To determine the publication rate of abstracts presented at the American Diabetes Association Scientific Sessions and to evaluate the association between statistical significance of study results and subsequent publication. Research Design and Methods. We conducted a retrospective cohort study of abstracts presented at the 2018 American Diabetes Association Scientific Sessions. The primary exposure was study result category (statistically significant vs. non-statistically significant findings), and the primary outcome was publication in an indexed journal within 5 years after conference presentation. Publication status was determined through PubMed/MEDLINE and Scopus searches. Adjusted relative risks (RRs) and 95% CIs were estimated using generalized linear models with Poisson distribution and robust variance. Results. Among 541 included abstracts, 321 (59.3%) were subsequently published in indexed journals. Abstracts reporting statistically significant findings had a higher publication rate than those reporting non-statistically significant findings (61.9% vs. 42.3%; p=0.002). In the adjusted analysis, abstracts with non-statistically significant findings had a lower likelihood of publication compared with those reporting statistically significant findings (adjusted RR 0.71 [95% CI 0.55-0.93]; p=0.013). Conclusions. Approximately four in ten abstracts presented at the ADA Scientific Sessions were not published within 5 years. Abstracts reporting non-statistically significant findings had a lower likelihood of subsequent publication, suggesting persistent publication bias in diabetology research. Future initiatives promoting the interpretation of effect estimates, confidence intervals and clinical relevance, rather than statistical significance alone, may help reduce selective dissemination of evidence
Ezeanosike, O. B.; Ezeanosike, E.; Anoke, C. I.; Okoro, O.; Orjingene, O.; Chukwu, E.; Okoli, U.
Show abstract
Background. Nigeria carries one of the world's largest burdens of neonatal death and remains far from the Sustainable Development Goal target. Whether health financing and macroeconomic instability are associated with newborn survival has rarely been examined for neonatal mortality specifically. Methods. We conducted an ecological time-series analysis of national annual data, covering 1990-2024 for macroeconomic models (n = 35) and 2000-2023 for health-financing models (n = 24), the periods for which published data exist; no values were imputed. Neonatal mortality came from the UN Inter-agency Group for Child Mortality Estimation 2025 round with 90% uncertainty intervals, and other series from the World Development Indicators. The primary model regressed log neonatal mortality on government health expenditure per capita (purchasing power parity), out-of-pocket share and currency instability, with a linear trend, a post-break trend spline and Newey-West standard errors; first differences without trend terms were the main sensitivity analysis. The break was located by segmented regression; currency instability was tested under four constructions. Results. The decline broke around 2010, the trend moving from -0.74 to +0.14 deaths per 1,000 annually (F = 145.4, p < 0.001). The subsequent rise fell within estimation uncertainty (2012: 37.6, 90% interval 33.9-41.5; 2022: 39.3, 33.4-46.4), supporting stagnation rather than reversal; Demographic and Health Surveys concur, reporting 42 per 1,000 for the five years preceding the 1990 survey and 41 preceding the 2024 survey. Government health expenditure per capita was inversely associated with neonatal mortality (-0.040, 95% CI -0.051 to -0.029, p < 0.001; first differences -0.016, p = 0.033) and was the only expenditure measure surviving both specifications; share-of-GDP measures did not (p = 0.196 and 0.889) and correlated positively in raw terms. Currency instability showed no association under any construction (p = 0.65-0.83). Public expenditure per capita moved non-monotonically, peaking in 2005, falling by 2010 and recovering by 2023 to a level still below the 2005 peak. Conclusions. Neonatal mortality in Nigeria is ecologically associated with public health expenditure per capita, but not with commonly used share-based measures, nor with currency instability. Rising public spending accompanied stalled progress, directing attention toward how health resources are converted into services. Annual modelled mortality estimates could not support year-to-year inference, a limitation relevant to comparable studies
Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.
Show abstract
Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.
Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.
Show abstract
Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Ricarte Almeida, E. R.; Mata Quintero, C. J.; Sesma Chazaro, J.; Peralta Rivera, C.; Arteaga Gonzalez, C. D.
Show abstract
Background: Sleeve gastrectomy is the most frequently performed bariatric procedure worldwide but is associated with the development of de novo gastroesophageal reflux disease (GERD). Hiatal hernia has been identified as a relevant anatomical factor in postoperative reflux, although most studies evaluate it dichotomously without analyzing whether its size influences GERD risk. The aim was to evaluate the association between preoperative hiatal hernia size and de novo GERD after sleeve gastrectomy. Methods: Retrospective, single - center, observational study of patients undergoing sleeve gastrectomy at Hospital Central Norte de Petroleos Mexicanos (2018 - 2025). Demographic and clinical characteristics, endoscopic classification of hiatal hernia size (small <2 cm, medium 2.1 - 4 cm, large >4 cm), and evidence of de novo GERD were analyzed using descriptive statistics, Fisher's exact test, odds ratio (=R) estimation with 95% confidence intervals (CI), and binary logistic regression. Statistical significance was set at p<0.05. Results: Fiftysix patients were included (mean age 48.3 {+/-} 8.1 years; 67.9% male). Hiatal hernia classification was conclusive in 46 patients (82.1%): 63.0% no hernia, 4.3% small, 30.4% medium, and 2.2% large. De novo GERD occurred in 14.0% of patients without preexisting GERD (6/43). No significant association was found between hiatal hernia size and de novo GERD (Fisher p=0.515). In the reduced logistic model, neither hiatal hernia (medium/large vs. absent/small; OR 3.47; 95% CI 0.50 - 29.43; p=0.207) nor age (OR 1.02; 95% CI 0.90 - 1.13; p=0.754) was significantly associated. No evaluated factor (sex, smoking, alcohol, age) reached significance. Conclusions: In this cohort, no statistically significant association was demonstrated between preoperative hiatal hernia size and de novo GERD after sleeve gastrectomy; however, the low number of events limits the ability to exclude a clinically relevant association. These findings are compatible with a multifactorial mechanism rather than with the isolated presence of this finding. Prospective studies with larger sample sizes and standardized reflux assessment instruments are required to confirm these results.